Informatics in Medicine Unlocked
○ Elsevier BV
Preprints posted in the last 30 days, ranked by how well they match Informatics in Medicine Unlocked's content profile, based on 22 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Smid, J.; Jezdik, P.; Kalina, A.; Kudr, M.; Janca, R.
Show abstract
Background: Precise localisation of intracranial electrode contacts is essential for the interpretation of stereoelectroencephalography recordings and planning epilepsy surgery. In current clinical practice, this is typically a manual process, which is time-consuming and prone to variability. Existing automated solutions are often fragmented across multiple tools requiring technical expertise, limiting their adoption in routine clinical workflows. This study presents an open-source extension for 3D Slicer that provides an integrated, user-friendly standalone solution for the direct automatic detection of electrode contacts within a widely used medical imaging platform. Results: The proposed method combines anchor bolt-based initialisation, probabilistic segmentation of electrode structures, and non-linear modelling to precisely track true electrode trajectories. The approach was evaluated on a dataset comprising 78 cases from 73 patients, including 1,078 electrodes with 14,480 contacts. The method achieved high localisation accuracy, with a median (interquartile range) deviation of 0.10 (0.06, 0.15) mm. Only 7/1078 (0.65%) electrodes required manual correction; these specific cases were handled using tools provided within the proposed extension. Conclusions: The presented extension enables fast, accurate, and reproducible electrode contact localisation within a single integrated environment. By combining automation with intuitive user interaction, it significantly reduces processing time while maintaining clinical reliability. The tool's free availability as an extension in 3D Slicer lowers the barrier to adoption and supports the standardisation of workflows across clinical and research centres.
Razmjooei, F.; Ashayeri, H.; Jafarzadeh, Z.; Dabbaghabdollahi, P.; Jafarizadeh, A.
Show abstract
Background: Uveal melanoma (UM) and cutaneous melanoma (CM) both originate from the same cell line. This proposes the possibility of a shared mechanism between entities, requiring explicit investigation. Methods: Data from GWAS Catalog and DisGeNET were used to identify shared variation-disease associations (VDAs) between UM and CM. The results were validated using the Ensembl database. In the next step, the STRING database was used to identify the protein-protein interaction. Results: Subsequently, 109 unique VDAs were identified for UM and 880 for CM. However, only 2 VDAs were found to be shared among UM and CM in different ethnic groups. These shared VDAs were rs12203592 of the IRF4 gene, rs12913832 of the HECT and RLD domain-containing E3 ubiquitin protein ligase 2 (HERC2) gene. Notably, PPI network assessment through STRING showcased that OCA2 and IRF4 directly interacted with HERC2. Conclusion: While HERC2 acts as a poor prognostic factor in uveal melanoma, IRF4 status is a key prognostic indicator in both UM and CM. Identifying IRF4 allele contributions enables a better understanding of melanoma pathogenesis and fosters the development of disease-specific approaches.
Zaitsev, V.; Wei, C.-S.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWElectroencephalography (EEG) is a promising tool for automated detection of mild cognitive impairment (MCI) and dementia, but comparisons across studies are limited by inconsistent datasets and evaluation protocols. This study benchmarks ten deep learning models across four resting-state EEG datasets and eight binary classification tasks using a unified preprocessing pipeline and five-fold subject-wise cross-validation. Each experiment was repeated ten times. SCCNet obtained the highest mean subject-level accuracy, sensitivity, and F1 score, while ShallowConvNet achieved the highest mean segment-level accuracy, specificity, and precision. Subject-level aggregation improved mean accuracy for all evaluated models, and performance varied substantially across datasets and diagnostic tasks. Higher computational cost did not consistently correspond to better classification performance, with several compact architectures remaining competitive with substantially larger models. The results provide a reproducible reference for comparing EEG-based dementia classification models under consistent subject-independent evaluation conditions.
Motta, J. A.; Motta, M. d. M.; Fernandez, C.
Show abstract
In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide sequences (amino acid sequences) governed by a well-defined grammar, which allowed us to apply NLP techniques, specifically Part of Speech tagging (POS tagging). Our predictive model was built by combining two techniques: CRF (from the Markov model family), which performs the sequencing, and BiLSTM (a deep learning model) which captures the past and future content of the sequences. The training space was created with the sequences of 105 genes associated with approximately 27,000 pathogenic variants. The model was evaluated using the metrics precision, P-R and ROC curves, AUC, and confusion matrices. Its performance was also compared against five known methods for predicting pathogenic variants. The results show exceptional performance that exceeds expectations and places this new method at the state of the art for predicting pathogenic DNA sequences.
Yuan, W.; Wang, Z.; Wu, Q.; He, X.; Tan, J.; Wei, X.; Li, R.; Yin, Y.; Wang, D.; Wang, G.; Chen, T.
Show abstract
Objectives: To develop and externally validate a wall-focused deep learning framework for identifying composite unstable intracranial aneurysm phenotypes on dual-phase high-resolution vessel wall imaging (HR-VWI), and to visualize model attention on the aneurysm wall surface. Methods: This retrospective multicenter study included patients with intracranial aneurysms who underwent both non-contrast and contrast-enhanced HR-VWI. Center 1 was used for model development and patient-level five-fold out-of-fold assessment, whereas Centers 2 and 3 served as independent external validation cohorts. For each aneurysm, dual-phase local wall patches and larger spatial context patches were generated. The Wall-Constrained Encoding Network (WCE-Net) extracted mask-constrained local wall features, and a transfer-learning U-Net with Nested Transformers (UNesT) branch extracted spatial context information. Branch outputs were fused by logit-level stacking. Model performance was evaluated using discrimination, calibration, and decision curve analysis. Three-dimensional gradient-weighted class activation mapping (Grad-CAM) responses were projected onto the reconstructed aneurysm wall surface and compared with HR-VWI surface signal intensity. Results: A total of 629 patients with 773 aneurysms were included. The final fusion model achieved areas under the receiver operating characteristic curves (AUCs) of 0.908, 0.857, and 0.855 in Center 1, external Center 2, and external Center 3, respectively. Corresponding Brier scores were 0.119, 0.153, and 0.150. Surface Grad-CAM showed partial spatial overlap between model-attention hotspots and high-signal HR-VWI regions. Conclusions: Dual-phase wall-focused local-context fusion showed feasibility for identifying composite unstable intracranial aneurysm phenotypes across centers. Surface Grad-CAM provided anatomically referenced visualization of model attention.
Mohammad, U.; Parani, P.; Saeed, F.
Show abstract
Background and Objective Epileptic seizure prediction is a critical challenge requiring the discrimination of subtle preictal physiological changes from interictal brain activity. While deep learning has shown promise in this domain, existing models often face limitations due to small EEG datasets, high computational costs for training from scratch, and a lack of patient-independent generalizability. In this paper, we present a novel framework for EEG-based seizure prediction that leverages pre-trained Vision Transformers (ViTs) through custom architectural modifications and optimized re-training strategies. Methods Our primary contributions include: [bullet]CVIT-ESP: A family of vision transformer architectures that replaces standard patch embedding layers with custom N-dimensional CNN stages to refine EEG representations. [bullet] ESPFormer: A lightweight, custom-designed transformer specifically engineered to mitigate overfitting on limited-scale EEG datasets. We identified optimal fine-tuning combinations for transformer blocks by devising a heuristic search-space reduction strategy, significantly reducing the training complexity. We validated our methods using the patient-independent MLSPred-Bench, involving 12 diverse benchmarks with varying seizure prediction horizons. Results Results demonstrate a clear progression in performance: while prior ResNet and vanilla Transformer models achieved an AUC-ROC of 69.0%, our CVIT-ESP architectures achieved the highest performance with a maximum average AUC of 76.4%. Conclusions These findings suggest that adapting pre-trained ViTs with domain-specific CNN front-ends and strategic fine-tuning offers a robust, generalizable, and resource-efficient path forward for clinical seizure prediction systems. Our code is available at: https://github.com/pcdslab/CVitEsp and https://github.com/pcdslab/ESPFormer
Bai, X.; Kishimoto, K.; Sugiyama, O.; TAMURA, H.
Show abstract
This study aims to improve the detection performance of age-related macular degeneration (AMD) in low-quality retinal images. BackgroundAMD is a leading cause of vision loss among older adults globally, and accurate detection is crucial for clinical management. However, low-quality optical coherence tomography (OCT) images significantly compromise diagnostic accuracy. ObjectiveTo enhance AMD detection in low-quality images using noise-augmented data augmentation and an improved YOLO deep learning model. MethodsPublic datasets from UCSD and Duke University were utilized; the training dataset comprised 24,980 OCT images (high-quality and noise-augmented low-quality), while the testing dataset included 1,000 images (584 AMD, 416 normal). The model is based on the YOLOv8n framework, integrated with Squeeze-and-Excitation blocks (SEblock) and Adaptive Sparse Self-Attention (ASSA), with an additional 160x160 detection layer for detecting small lesions. Evaluation metrics included accuracy, sensitivity, specificity, and F2-score. ResultsThe proposed model achieved an accuracy of 99.02%, sensitivity of 98.17%, specificity of 100%, and an F2-score of 98.50% on the Duke dataset. Detection rates were significantly improved compared to traditional methods, particularly in low-quality images, with a detection rate of 89.60%, markedly superior to original YOLOv8n (55.10%) and classical models like ResNet50. ConclusionThe enhanced model, employing noise-augmented training data and improved attention mechanisms, demonstrates excellent AMD detection capabilities in low-quality OCT images, showing broad potential for clinical applications.
Niessen, S.; Focke, C.; Keller, S.; Scheffold, H.; Hempel, S.; Lettner, J. D.; Scheef, T.; Klar, R. F. U.; Vladimirov, G.; Crossley, K. A.; Bittner, D.; Deuter, M.; Kissel, S.; Chikhladze, S.; Fichtner-Feigl, S.; Duyster, J.; Boerries, M.; Neubauer, J.; Scherer, F.; Luebbert, M.; Quante, M.; Ruess, D. A.; Becker, H.
Show abstract
Background Therapy resistance in pancreatic ductal adenocarcinoma (PDAC) is facilitated by the desmoplastic tumor microenvironment (TME) orchestrated by cancer associated fibroblasts (CAFs). Upon activation, pancreatic stellate cells (PSCs) deplete their intracellular retinoic acid (RA)-containing lipid droplets and secrete stromal remodeling proteins like pentraxin 3 (PTX3), leading to cancer progression. Preclinical evidence indicates that all-trans RA (ATRA) reprograms the TME, while circulating vitamin A and PTX3 were proposed as biomarkers for ATRA response in PDAC. To support further clinical development of RA-based therapies in PDAC, we studied the effects of ATRA on CAFs and patient-derived organoids (PDO) and evaluated the clinical relevance of these biomarkers in PDAC patients. Methods We employed viability assays in human and murine organoid mono- and co-culture models to explore the efficacy of adding ATRA to gemcitabine (GEM). In parallel, we conducted a prospective observational study and assessed vitamin A and PTX3 as response biomarkers in peripheral blood collected before first treatment and at cycles 2 and 4 of treatment among patients with advanced PDAC receiving GEM with or without nab-paclitaxel (NAB-P). Results In PDO monocultures, a significant additive effect of ATRA in combination with GEM on viability was observed in 5 (41%) of 12 PDOs and this effect was numerically more frequent in organoids from patients who had clinically responded to GEM. In human and murine 3D PDO+PSC/CAF co-cultures, ATRA demonstrated an additional direct impact on the viability of stromal cells. Clinically, among 18 patients with PDAC treated with GEM+/-NAB-P, patients with no treatment response (n=10) showed an increase in PTX3 and concomitant decrease in vitamin A levels under therapy. In contrast, response was associated with stable vitamin A levels and a trend towards lower PTX3 levels during chemotherapy. Conclusions Our preclinical data support the repurposing of ATRA, an agent with favorable toxicity profile, to potentiate the efficacy of GEM in PDAC treatment. Complementing these results, our clinical data suggest vitamin A and PTX3 as promising response biomarkers in PDAC treatment, not restricted to ATRA containing regimens.
Irajizad, E.; Lopez, C.; Chari, S.; Vykoukal, J.; Spencer, R.; Li, Y.; Dennison, J.; Koay, E.; McAllister, F.; Kim, M.; Young, M.; Hart, P.; Fischer, W.; Vandeneeden, S.; Wu, B.; Feng, Z.; Hanash, S.; Maitra, A.; Fahrmann, J.; Consortium for the Study of Chronic Pancreatitis, Diabetes, and Pancreatic Cancer (CPDPC),
Show abstract
PURPOSE: To assess the predictive performance of panel protein biomarkers as well as an established algorithm that considers repeat biomarker testing for risk prediction of PDAC among a prospective cohort of patients with New-onset diabetes. PATIENTS AND METHODS: A panel of protein biomarkers (CA19-9, CA125, CEA, LRG1, REG3A and TIMP1) were assayed in 6,516 serially collected pre-diagnostic plasma samples from 2,121 NOD patients from the Consortium of Chronic Pancreatitis Diabetes and Pancreatic Cancer (CPDPC)-initiated NOD study who completed the 3-year study follow-up period. The specimen set included 25 pre-diagnostic samples from the 12 PDAC cases diagnosed during study follow-up. We applied a single threshold (ST) method, which considers biomarker levels at a single time point, as well as a previously established parametrical empirical Bayes (PEB) algorithm, which considers prior biomarker measurements, with case calls made based on pre-specified cutoffs corresponding to 1% 1-year risk. Resultant biomarker data as well as case calls were provided to the EDRN Data Management and Coordinating Center as part of a Prospective-sample-collection-Retrospective-Blinded-Evaluation (ProBE)-compliant Phase 3 biomarker validation study. Area under the Receiver Operating Characteristic Curves (AUC), sensitivity, specificity, population-level positive predictive value (PPV), and negative predictive value (NPV) are reported. RESULTS: The 3-year incidence of PDAC in the NOD cohort was 0.57%. When considering PDAC vs non-cancer controls, respective AUCs of individual protein biomarkers ranged from 0.52-0.94, with CA19-9 achieving the highest overall performance of 0.94 (95% CI: 0.86-1.00). At the pre-defined 1% 1-year risk threshold, CA19-9 yielded sensitivity of 83.3% at 97.2% specificity. Additional markers CEA, CA125, and TIMP1 demonstrated sensitivity of 33.3%, 41.7%, and 8.3%, respectively. In a subset of patients, CA19-9 first tested positive at a median (interquartile range [IQR]) of 7 months (4 to 14 months) prior to clinical PDAC diagnosis. Of the two PDAC cases missed by CA19-9 using the ST method, one (diagnosed with stage III PDAC) was detected using the PEBCA19-9 algorithm. CONCLUSION: In the setting of adult new onset diabetes, CA19-9 is a readily available and promising biomarker that can be leveraged for earlier detection of an underlying pancreatic cancer. Additional protein biomarkers may improve sensitivity for earlier detection of PDAC among cases with low CA19-9.
Kwon, S.; Lee, C. S.; Lee, A. Y.; Zhang, L.
Show abstract
Purpose: To evaluate whether fluorescence lifetime imaging ophthalmoscopy (FLIO) combined with deep learning can detect metabolic signatures for classification of type 2 diabetes mellitus (T2DM). Design: Cross-sectional analysis of participants included AI-READI dataset (version 3) with FLIO imaging and and hemoglobin A1c (HbA1c) measurement. Subjects: 1,783 participants from the AI-READI dataset (version 3) with HbA1c measurements and FLIO imaging scans (6,912 total): 671 normoglycemic, 726 prediabetic, and 386 diabetic. Methods: Mean fluorescence lifetime maps were generated using a center-of-mass approach and used as inputs to AI models. We trained convolutional neural networks (CNNs), ResNet-18, and XGBoost under three-class (normal, prediabetic, diabetic) and two binary (normal vs. impaired; normal vs. diabetic) classification schemes, using nested 5-fold cross-validation with participant-level grouping. Main Outcome Measures: Macro-averaged accuracy, F1 score, area under the receiver operating characteristic curve (AUROC), sensitivity, specificity, and positive predictive value (PPV). Results: Group-averaged lifetime maps demonstrated consistent spatial differences across glycemic groups, with progressively longer lifetimes from normal to diabetic participants. The CNN achieved the best overall performance in the 3-class classification (accuracy 0.41 +/- 0.03, F1 score 0.39 +/- 0.02, AUROC 0.58 +/- 0.02), compared to the random classifier for 3-class classification (AUROC = 0.50; accuracy = F1 = 0.33). ResNet-18 and XGBoost showed similar performance (AUROC 0.53-0.58). Confusion matrices revealed substantial overlap between classes, with frequent misclassification toward the prediabetes group. Binary reformulation (normal vs. diabetic) improved performance substantially, with the CNN resulting in AUROC 0.63 +/- 0.02 and XGBoost 0.67 +/- 0.07. Conclusions: FLIO-derived lifetime maps capture metabolic signals associated with glycemic status but yield modest classification performance with current AI models. These findings highlight both the potential and the challenges of using FLIO for early metabolic screening and monitoring, informing future development of clinically applicable imaging biomarkers.
Khandelwal, S.; Jarvis, N.; Zhan, J.
Show abstract
Glioblastoma (GBM) is a highly aggressive brain tumor with an extremely poor 5-year survival rate of 6.9%, largely attributable to the lack of reliable biomarkers. While competing endogenous RNA (ceRNA) and copy number variation (CNV) analyses offer unique biomarker identification potential, current approaches neglect the integration of multiple regulatory mechanisms for biomarker detection. To address this limitation, we applied relational graph convolutional networks (RGCNs) to ceRNA and CNV knowledge graphs through a novel late fusion ensemble architecture. The proposed architecture outperformed baseline models and identified five novel biomarkers, including hsa-miR-196a and hsa-miR-224. Kaplan-Meier survival analysis and Cox regression indicated that the identified genes hold significant prognostic and diagnostic power. The early stratification of the Kaplan-Meier curves indicates the potential these genes hold for patient survival prediction. The results illustrate that a late fusion RGCN ensemble effectively captures complex gene interactions, overcoming limitations of existing models and providing a framework for biomarker discovery. The novel biomarkers serve as prospective targets for future GBM therapeutic development and candidates for non-invasive diagnostic assays.
Płonka, W.; Kostka, D.; Lalik, A.; Kurpas, M.; Dinh, K. N.; Sitkiewicz, M.; Kimmel, M.; Rzyman, W.; Jaksik, R.
Show abstract
Formalin-fixed, paraffin-embedded (FFPE) tissues remain an essential resource for molecular studies, yet formalin-induced cytosine deamination introduces characteristic C>T/G>A artifacts that compromise the accuracy of next-generation sequencing (NGS) analyses. Numerous computational methods and enzymatic DNA repair strategies have been proposed to reduce these artifacts, but no systematic comparison across tools and experimental conditions exists. Here, we evaluate the performance of seven computational approaches (SOBDetector, Ideafix, MicroSEC, FFPolish, DeepOmics FFPE/FFPE-PLUS, FFPErase) together with the NEBNext(R) FFPE DNA Repair Mix v2, a multi-enzyme repair system applied during DNA preparation. Using three independent datasets, one based on whole genome sequencing (CGCI-BL) and two on whole exome sequencing (TCGA-PC and SUT-LUAD, the latter containing enzymatically repaired samples), and matched fresh-frozen samples as the gold standard, we assess precision, sensitivity, and artifact reduction efficiency across all methods. We further examine the potential synergy between enzymatic repair and post-sequencing computational filtering. Our results provide practical guidelines for FFPE artifact correction and demonstrate that enzymatic treatment provides the best results, while among the computational methods, FFPErase offers the most robust reduction of cytosine deamination artifacts while maximizing the retention of true somatic variants. KEY MESSAGESO_LIFormalin fixation in FFPE samples introduces artifacts that can significantly affect the accuracy of NGS analyses. C_LIO_LIAmong the evaluated approaches, enzymatic repair using NEBNext(R) FFPE DNA Repair Mix v2 achieves the most effective reduction of sequencing artifacts. C_LIO_LIComputational methods vary in performance, with FFPErase showing the most robust balance between artifact removal and retention of true somatic variants. C_LIO_LICombining enzymatic repair with computational filtering did not lead to consistent improvements in performance across datasets. C_LI
Lu, Z.; Uddin, S.; Uribe, S.; White, S.; Martins, R. T.; Chau, S.; Mosaddek, A. S. M.; Islam, M. S.; Nahar, N.; Azad, A. K. M.; Hossain, K. M. N.; Choudhury, H. S.; Hasan, K. M. R.; Mosaddek, N.; Rahman, S.; Hossain, M. M.; Sizar, K. M. M. H.; Angione, C.; Lio, P.; Islam, M. T.; Moni, M. A.
Show abstract
Stroke remains a leading cause of mortality and long-term disability worldwide, yet rapid diagnosis is often limited by the shortage of trained radiologists, particularly in resource-constrained settings. Automated analysis of CT imaging offers a potential solution, but existing methods often struggle to achieve clinically generalisable performance while jointly addressing multiple diagnostic tasks. Here we present the Intelligent Integrated Stroke Diagnosis System IISDS, an end-to-end deep learning framework built upon StrokeGNN, a graph-based architecture that integrates 3D contextual feature extraction with U-Net-based 2D lesion segmentation to enable comprehensive stroke analysis from non-contrast CT scans. IISDS performs stroke subtype classification, lesion segmentation and lesion volume estimation within a unified pipeline. To develop and validate the system, we collected and curated BGD-ISD through a collaboration between AI researchers, neurologists, radiologists and clinicians, resulting in a large multi-centre dataset comprising 1,507 CT scans from 597 stroke cases acquired across six hospitals and medical centres in Bangladesh. Across BGD-ISD and multiple publicly available datasets, IISDS achieves state-of-the-art performance on all tasks, improving segmentation accuracy by [≥]0.011 Dice score, reducing lesion volume estimation error by [≥]0.3 average symmetric surface distance (ASSD), and increasing classification performance by [≥]0.018 area under the receiver operating characteristic curve (AUC) compared with existing approaches. These results demonstrate the potential of graph-based deep learning to enable clinically generalisable, automated and scalable stroke diagnosis from CT imaging, supporting rapid clinical decision-making, particularly in healthcare environments with limited access to expert radiological interpretation.
Quan, W.; Henault, D.; Zhang, A.; Jang, G. H.; Hasnain, S. M.; Bevacqua, D.; Deng, Y.; Flores-Figueroa, E.; Ni, K.; Light, N.; Wilson, J. M.; Dodd, A.; Tsang, E. S.; King, D. A.; Habowski, A. N.; Yu, K.; Perez, K.; Aguirre, A. J.; O'Reilly, E. M.; Wolpin, B. M.; Pugh, T. J.; Tuveson, D. A.; Jaffee, E. M.; Gallinger, S.; O'Kane, G.; Notta, F.; Knox, J. J.; Grant, R. C.
Show abstract
Purpose Modified FOLFIRINOX (FFX) and gemcitabine plus nab-paclitaxel (GNP) are standard first-line treatments for metastatic pancreatic ductal adenocarcinoma (PDAC), but no validated biomarker guides treatment selection. We developed MULTIPL, a multimodal machine learning system, and established the PASS-01 Challenge to benchmark prognostic and predictive biomarkers. Patients and Methods MULTIPL was trained in the COMPASS study (N=268), integrating clinical, digitized histopathology, whole-genome, and RNA-seq data. MULTIPL, PurIST, hENT1 expression, and HRDetect were evaluated in the PASS-01 trial, a randomized phase II trial of FFX versus GNP (N=160), within the Challenge. The primary endpoint was differential treatment benefit measured by concordance-for-benefit for progression-free survival. Results MULTIPL had the highest concordance index for OS among individually evaluated biomarkers (0.595; 95% confidence interval [CI], 0.55-0.65) and separated high- versus low-risk patients (hazard ratio, 1.62; 95% CI, 1.13-2.33; P=0.009). Patients recommended for GNP by MULTIPL had significantly longer OS with GNP than with FFX (hazard ratio, 0.47; 95% CI, 0.28-0.82; P=0.007), whereas patients recommended for FFX had similar OS between treatments. Interpretability analysis of MULTIPL in COMPASS identified KDM6A alterations and SSTR1 expression as prognostic biomarkers, which were validated in PASS-01. However, none of the tested biomarkers significantly predicted differential treatment benefit in the PASS-01 Challenge. Conclusion MULTIPL demonstrated robust prognostic performance in external validation, identified a subgroup enriched for benefit from GNP, and enabled discovery and validation of prognostic biomarkers in metastatic PDAC. However, no biomarker met the primary endpoint for differential treatment benefit, underscoring the value of the PASS-01 Challenge.
Gomez Bergna, S. M.; Amoros Morales, L. C.; Gonzalez Abad, A.; Vilches, J.; Tongiani, S. E.; Salvador, R.; Romanowski, V.; Pidre, M. L.; Ferrelli, M. L.
Show abstract
Spodoptera frugiperda is one of the most important agronomical pests due to its migratory capacity and broad host range. Since it is resistant to several insecticides, novel control strategies are being explored to control it. In this way, Spodoptera frugiperda Multiple Nucleopolyhedrovirus, a natural pathogen, has been proposed for its biocontrol. In this work, we performed a small RNA-seq on uninfected larvae and larvae infected with SfMNPV to identify expressed miRNA, characterize them, and identify differentially expressed (DE) miRNA in the infected condition. We identified several known and putative novel miRNAs, some of which are encoded in multiple copies and may be expressed within miRNA clusters. We also found 13 DE miRNA, most of them previously reported, two of them are putative novel miRNAs identified in this work. We predicted miRNA targets and found that their putative biological role could be related with processes relevant to the infection such as proliferative and apoptotic pathways, cell cycle regulation, autophagy, DNA damage response (DDR), vesicle transport, cytoskeleton remodelling, JAK/STAT and Toll signaling pathway, and immune response activation, among others. Moreover, we observed that several of the putative targets were hub genes in a predicted protein - protein interaction network. Finally, we found DE miRNA putatively associated with the regulation of viral gene expression, suggesting they might have a role in modulating the infection. Our results contribute to better understanding the miRNA landscape in S. frugiperda, and their putative role upon SfMNPV infection.
Udumanne, T. P.; Liew, Y. J.; Pascovici, D.; Yang, T.; Lee-Ng, K. K. M.; Gracie, G.; Kumarasinghe, P.; McLeod, D.; Brown, I.; Bourke, M. J.; Lord, S. J.; Ross, J.; Lord, R. V.
Show abstract
Esophageal adenocarcinoma (EAC) has a poor five-year survival rate and one of the fastest-rising incidences of any cancer. The presence of dysplasia in Barrett's esophagus (BE) is the main risk factor for EAC development and guides clinical management. Unfortunately, the current histopathological diagnosis of dysplasia is unreliable, with poor inter-observer agreement, highlighting the need for novel biomarkers that can improve diagnostic accuracy. Here, we performed transcriptome profiling across the full spectrum of BE-related neoplasia in 85 samples to delineate gene expression alterations in progressively worse disease stages and identify biomarkers that could complement histopathology to improve the detection of dysplasia and EAC in endoscopic biopsy specimens. Differential gene expression and pathway analyses revealed that the most extensive transcriptional changes occurred during the transition from normal squamous (NSq) to non-dysplastic BE (NDBE), consistent with metaplastic transformation. Compared to NDBE, dysplasia was characterized by enhanced cellular growth and proliferation; upregulation of immune processes and oncogenic signaling pathways were present in EAC. Using machine learning approaches, we identified a novel five-gene panel suitable for a potential RNAseq-based diagnostic test (SLC11A1, IL36A, LUCAT1, MIR215, RNU6-954P) and performed an initial validation of this signature in an additional 51 samples. We also identified several potential novel immunohistochemical markers that may warrant further evaluation, including TREM1, CXCL5, OSM, and motilin. In summary, by delineating transcriptional changes across the full disease spectrum, this study identifies several candidate biomarkers for improving current diagnostic methods for Barrett's dysplasia and EAC.
Rounds, C. C.; Ravi, D.; Huang, G.; Mengesha, B.; Tran, S.; Garcia, A.; Rueb, N.; Chang, Y. H.; Park, B. S.; Wong, M. H.; Gibbs, S. L.
Show abstract
SignificanceRare-cell identification in fluorescence microscopy remains challenging because targets are sparse and background varies between specimens. Combining specimen-specific fluorescence enrichment with image classification may enable efficient and more specific automated detection of rare cells. AimWe developed a two-stage framework to identify and quantify candidate rare circulating hybrid neoplastic cells (CHCs, ECAD+/CD45+) in peripheral blood mononuclear cell (PBMC) preparations from tumor-bearing and tumor-naive mice. ApproachPBMCs from 28 mice were imaged by multichannel fluorescence microscopy. Matched unstained samples established animal-specific ECAD and CD45 background distributions for candidate cell enrichment. Blinded multi-annotator consensus labels were used to train a convolutional neural network (CNN) from DAPI, ECAD, and CD45 image crops. Generalization was evaluated by leave-one-animal-out validation across 10 random initializations. Final classification used a 10-model ensemble, and rare-cell burden was compared between groups using negative-binomial regression with total segmented-cell count as an exposure. ResultsOf the 1,065,512 segmented cells, enrichment retained 10,176 candidates (0.96%), reducing the search space by >99%. Four of five evaluable tumor-bearing animals showed reproducible held-out discrimination, with median quantified area under the receiver operator characteristic curve (AUROCs) of 0.918-0.951; one animal was a reproducible outlier (median AUROC, 0.338). Ensemble deployment identified 157.94 positive-consensus cells per 50,000 segmented cells in tumor-bearing animals versus 49.55 in controls. The estimated rare-cell rate was 3.15-fold higher in tumor-bearing animals (95% CI, 0.91-10.99; two-sided p=0.071; prespecified one-sided p=0.036). ConclusionsSpecimen-specific fluorescence enrichment combined with supervised image classification reduced the cellular search space and enabled automated quantification of a rare CHC (ECAD+/CD45+) phenotypes. Cross-animal validation also identified specimen-specific generalization failure, highlighting the importance of biological-specimen-level validation.
Feng, W.; Liu, S.; Yang, Z.; Tao, Y.; Gu, X.; Jin, W.
Show abstract
Background Hepatocellular carcinoma (HCC) treatment selection demands nuanced integration of heterogeneous patient data, yet prevailing predictive models rely on restricted data modalities and oversimplified therapeutic frameworks, compromising clinical translation. Objective We developed and validated a multimodal artificial intelligence framework to guide optimal treatment strategy selection across the full spectrum of HCC interventions. Methods This retrospective study comprised 1,043 HCC patients (development cohort, January 2017-December 2023) and 55 external validation patients (2023) from Wuxi Peoples Hospital. We engineered Embedding-Augmented Extra Trees (ET-Emb), a novel model fusing structured clinical variables with contextual text embeddings derived from medical histories and radiology reports. ET-Emb quantifies probabilities for five primary treatments: open/laparoscopic resection, transarterial chemoembolization, radiofrequency ablation (RFA), and chemotherapy. Model performance was rigorously assessed via 10-fold cross-validation and external validation using ROC-AUC and PR-AUC metrics. Results ET-Emb demonstrated robust performance in the development cohort (ROC-AUC: 0.84 {+/-} 0.04; PR-AUC: 0.55 {+/-} 0.06), significantly outperforming established benchmarks. This generalizability was preserved in external validation (ROC-AUC: 0.77 {+/-} 0.02; PR-AUC: 0.47 {+/-} 0.03). SHAP analysis identified textual clinical narratives and socioeconomic determinants as critical predictive drivers. Conclusions By unifying structured and unstructured data modalities, ET-Emb delivers accurate, multi-treatment strategy prediction for HCC. Its clinical validity and the demonstrated significance of textual features establish multimodal AI as an essential paradigm for simulating complex oncological decision-making, positioning ET-Emb as a transformative tool for precision HCC management.
Plabon, A. M.; Mukit, A.; Neyamul, M.; Jehady, O. F.; Zuba, F. T.; Mina, M. F.; Islam, T.
Show abstract
Interictal epileptiform discharges (IEDs) are diagnostically important EEG abnormalities observed between seizures. This study addresses a conditional spatial-classification task where every analyzed four-second epoch had already been reviewed and confirmed by experts as containing an IED, and the model assigned that epoch to one of five predefined scalp-distribution categories (generalized, frontal, temporal, occipital, or centro-parietal). The analysis therefore does not evaluate IED-versus-non-IED detection. After preprocessing, 2,514 IED-labelled epochs were analyzed using identical stratified epoch-level partitions, SMOTE based training, 26 handcrafted features per included channel, and multiple machine-learning classifiers. A staged channel ablation compared 19-channel scalp EEG, 21-channel EEG with ECG, and the complete 29-channel input containing scalp EEG, referential, ECG, and EMG channels. The best EEG-only result was obtained with linear discriminant analysis (88.89% test accuracy). CatBoost achieved 93.25% on EEG with ECG channel and 94.44% with the whole channel set. All eight directly comparable classifiers showed numerically higher test accuracy after ECG channel was added; for CatBoost, the increase was 6.35 percentage points. In the EEG with ECG channel, CatBoost model on ECG channel on right and left arm received respectively 15.79% and 15.12% of normalized global SHAP attribution, and beta-band power was the leading of all features (18.76%). These SHAP values indicate model-specific predictive contributions and do not establish physiological biomarkers, causal autonomic mechanisms, or clinical localization. The findings support a limited methodological conclusion which is ECG-derived features were associated with improved internal epoch-level categorization of expert-confirmed IED epochs. They do not establish IED detection, artifact rejection, independent EMG effects, or generalization to unseen patients.
Hassan, S.; Razaulla, S. M.; Pandey, R. K.
Show abstract
Group A Streptococcus (GAS), or Streptococcus pyogenes is almost exclusive and greatly adapted human pathogen. It causes a wide array of clinical symptoms, ranging from minor infections of the skin and soft tissues to pharyngitis, meningitis, pneumonia, bacteraemia, cellulitis, puerperal sepsis, and necrotising fasciitis. The risk of S. pyogenes infection is known to be influenced by several host characteristics, including age, underlying diseases like diabetes, varicella, or skin lesions, both chronic and acute, and certain risk behaviours such as use of drugs. Household size and overcrowding are two environmental factors that significantly affect the transmission of S. pyogenes. The majority of cases occur spontaneously in the community, and preventative opportunities are still limited. A large portion of GAS-related mortality is found in low-income areas and communities. Based on aforementioned public health risk, the creation of effective therapeutic vaccines would be an excellent addition to current control measures. The purpose of this work is to address the need for new instruments to aid in the elimination of S. pyogenes infections. The discovery of high antigenic regions in several highly conserved proteins brings us one step closer to developing peptide vaccines capable of influencing the different phases of S. pyogenes infection, providing more effective defence and greater serotype coverage. This study used various techniques of immunoinformatics to design an effective multi-epitope vaccine that produced neutralising antibodies against multiple strains of S. pyogenes.